Bulk Data and Subscriptions
The FHIR REST API is designed for one patient at a time. Two other mechanisms exist for the cases it does not serve: Bulk Data Access for population-scale extraction, and Subscriptions for event notification.
Choosing between them — and knowing when neither is the answer — is a recurring architecture decision in analytics, surveillance and exchange design.
Bulk Data Access
- Tier 1 · https://hl7.org/fhir/uv/bulkdata/
Also called Flat FHIR. It exports large volumes of FHIR resources as NDJSON — newline-delimited JSON, one resource per line — via an asynchronous request-and-poll pattern.
The flow
1. Client authenticates SMART backend services (client credentials + JWT)
│
2. Kick-off request
GET [base]/Group/[id]/$export?_type=Patient,Observation&_since=2026-01-01
Prefer: respond-async
│
3. Server responds 202 Accepted, Content-Location: <status URL>
│
4. Client polls the status URL 202 (in progress, with X-Progress) …
│ … 200 with a manifest when complete
5. Manifest lists NDJSON file URLs, per resource type
│
6. Client downloads files (often from object storage, with their own auth)
│
7. Client deletes the job (DELETE on the status URL)
Three export levels: $export on the system root (everything the client may
see), on Patient (all patients), and on Group (a defined cohort — the usual
choice, because it makes the authorisation question answerable).
Where it fits
- Loading a data warehouse or lakehouse from an EMR
- Quality measure computation and payer–provider reporting
- Research cohort extraction
- Migrating between systems
- Training and evaluating machine learning models
What to design for
| Concern | Reality |
|---|---|
| Volume | A national export is terabytes. Plan object storage, not a POST body. |
| Duration | Hours. The client must survive restarts and resume from the manifest. |
| Incremental | Use _since; a full export nightly does not scale. Confirm the server's _since semantics — "changed since" is not always what it means. |
| Authorisation | system/*.rs scopes grant a great deal. Scope by Group and log every export as an AuditEvent. |
| Governance | A bulk export is a bulk disclosure. It needs a legal basis, a purpose, a retention period and a named recipient — the technical ease of it is exactly the risk. |
| De-identification | The specification does not de-identify. If the destination is analytics, that is a separate, deliberate step — see health data. |
The most common failure is treating bulk export as an ETL convenience and
skipping the governance. A $export endpoint with a broad system scope is a
one-request copy of the national health record.
Subscriptions
FHIR's mechanism for "tell me when something happens" rather than polling.
Two generations
R4 Subscription — the client registers a search criteria string and a
channel (rest-hook, websocket, email, message). The server notifies on match.
Widely implemented, but the semantics of what is delivered are underspecified,
and it does not scale well to many subscribers.
R5 Subscriptions Backport (SubscriptionTopic) — the server publishes named
topics with defined triggers and payloads; clients subscribe to a topic
rather than inventing a query. This is the direction of travel and is available
for R4 via the backport implementation guide:
https://hl7.org/fhir/uv/subscriptions-backport/.
Prefer topic-based subscriptions for new work. Search-criteria subscriptions put the burden of correctness on every subscriber and make server-side optimisation nearly impossible.
Payload types
| Type | Delivers | Use when |
|---|---|---|
empty | "Something changed" only | The subscriber will re-query; safest for privacy |
id-only | The resource reference | The subscriber fetches with its own authorisation |
full-resource | The whole resource | Trusted subscriber, latency matters |
id-only is the usual right answer: the notification channel carries no
clinical data, and the fetch is separately authorised and audited.
Where it fits
- Notifiable disease reporting to surveillance
- Result-available notification to an ordering system
- Keeping a shared health record or index current
- Triggering decision support or an alerting workflow
What to design for
- Delivery is at-least-once, at best. Subscribers must be idempotent.
- Notifications can be lost. Pair every subscription with a periodic reconciliation query. A system whose correctness depends on never missing a webhook will be wrong eventually.
- Ordering is not guaranteed. Use resource versions and timestamps, not arrival order.
- Back-pressure. A campaign or a bulk import produces a notification storm; the subscriber's endpoint must be rate-limited and queued, not overwhelmed.
- Endpoint security. A rest-hook endpoint is an inbound path into the subscriber. Authenticate it, and verify the notification's origin.
- Lifecycle. Subscriptions go stale. Expire them, monitor error counts, and disable failing ones with an alert rather than retrying forever.
Choosing the mechanism
| Need | Mechanism |
|---|---|
| One patient's record, now | FHIR REST search / $everything |
| Everything, periodically, for analytics | Bulk Data $export |
| Tell me when this specific thing happens | Subscription (topic-based) |
| Reliable ordered stream between two internal systems | A message broker (Kafka, NATS) behind the interoperability layer — not FHIR Subscriptions |
| Transactional multi-resource write | FHIR transaction Bundle |
| Legacy hospital event feed | HL7 v2 messaging |
Subscriptions are a notification mechanism, not an event bus. Using them as the backbone of an event-driven architecture leads to the failure modes above at a scale that is hard to recover from — see event-driven interoperability.
References
- FHIR Bulk Data Access IG — https://hl7.org/fhir/uv/bulkdata/
- FHIR Subscriptions Backport IG — https://hl7.org/fhir/uv/subscriptions-backport/
- FHIR R5 Subscription — https://hl7.org/fhir/subscription.html
- SMART backend services — https://hl7.org/fhir/smart-app-launch/backend-services.html
- Asynchronous request pattern — https://hl7.org/fhir/async.html